Papers with dialogue quality

17 papers
CPO: Addressing Reward Ambiguity in Role-playing Dialogue via Comparative Policy Optimization (2025.findings-emnlp)

Copied to clipboard

Challenge: Comparative Policy Optimization (CPO) redefines the reward evaluation paradigm by shifting from sample-wise scoring to comparative group-wise score.
Approach: They propose a method to optimize subjective tasks by shifting from sample-wise to comparative group-wise scoring.
Outcome: The proposed framework shifts from sample-wise scoring to comparative group-wise score . it minimizes contextual bias and enables more robust and fair performance evaluation.
SelF-Eval: Self-supervised Fine-grained Dialogue Evaluation (2022.coling-1)

Copied to clipboard

Challenge: Existing evaluation metrics are expensive and easy to conduct but ineffective to reflect dialogue quality.
Approach: They propose a self-supervised fine-grained dialogue evaluation framework which can automatically assign fine-granular scores for arbitrarily dialogue data.
Outcome: The proposed framework is highly consistent with human evaluations and better than the state-of-the-art models.
Bridging Cultural Nuances in Dialogue Agents through Cultural Value Surveys (2024.findings-eacl)

Copied to clipboard

Challenge: integrating cultural dimensions with dialogue encoding features can enhance the predictive accuracy and quality of dialogue agents.
Approach: They propose to incorporate cultural dimensions into dialogue encoding features to enhance the predictive accuracy of dialogue agents.
Outcome: The proposed model improves the accuracy and quality of dialogue predictions by incorporating cultural dimensions with dialogue encoding features.
Filtering Noisy Dialogue Corpora by Connectivity and Content Relatedness (2020.emnlp-main)

Copied to clipboard

Challenge: Large-scale dialogue datasets contain a non-negligible number of unacceptable utterance pairs . previous studies have identified such flaws and reported that the corpus is noisy .
Approach: They propose a method for scoring the quality of utterance pairs based on their connectivity and relatedness.
Outcome: The proposed method has a good correlation with human judgment of dialogue quality and is applied to training data filtered by the proposed method.
AugESC: Dialogue Augmentation with Large Language Models for Emotional Support Conversation (2023.findings-acl)

Copied to clipboard

Challenge: Crowdsourced dialogue corpora are limited in scale and topic coverage due to the expensive cost of data curation.
Approach: They construct an augmented dataset for the emotional support conversation task using large language models for dialogue augmentation.
Outcome: The proposed approach outperforms baselines of dialogue augmentation and improves the model's generalization ability to open-domain topics.
DialogCC: An Automated Pipeline for Creating High-Quality Multi-Modal Dialogue Dataset (2024.naacl-long)

Copied to clipboard

Challenge: Existing multi-modal dialogue datasets that focus on image-based dialogues have low quality and limited diversity of images per dialogue.
Approach: They propose to construct a multi-modal dialogue dataset that guarantees both dialogue quality and image diversity without requiring minimum human effort.
Outcome: The proposed dataset outperforms existing datasets in terms of quality and diversity in human evaluation.
Building Persona Consistent Dialogue Agents with Offline Reinforcement Learning (2023.emnlp-main)

Copied to clipboard

Challenge: Existing methods to improve persona consistency are centered around supervised learning or online reinforcement learning (RL). Existing approaches to improve consistency are expensive and require additional training.
Approach: They propose an offline supervised learning framework to improve persona consistency of dialogue systems by punishing and rewarding specific utterances.
Outcome: The proposed framework improves both the persona consistency and dialogue quality of a state-of-the-art social chatbot.
Dream to Chat: Model-based Reinforcement Learning on Dialogues with User Belief Modeling (2025.findings-emnlp)

Copied to clipboard

Challenge: a framework for constructing dialogue world models for natural language tasks is currently lacking.
Approach: They propose a framework that can be used to train a dialogue world model.
Outcome: The proposed framework can predict future utterances and user beliefs . it can achieve state-of-the-art performance on emotion classification and sentiment identification .
MTR-DuplexBench: Towards a Comprehensive Evaluation of Multi-Round Conversations for Full-Duplex Speech Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks focus on evaluating single-round interactions, neglecting other critical aspects.
Approach: They propose a benchmark to evaluate full-duplex speech language models in multi-round settings . they segment continuous full-dual dialogues into discrete turns for evaluation .
Outcome: The proposed benchmark compared full-duplex speech language models with full-dual speech models . the results show that the models perform better in multi-round settings than standard models compared to benchmarks .
From Personas to Talks: Revisiting the Impact of Personas on LLM-Synthesized Emotional Support Conversations (2025.emnlp-main)

Copied to clipboard

Challenge: Experimental results show that LLMs can infer persona traits and subtle shifts in emotionality and extraversion occur . scalable solutions with reduced costs and enhanced data privacy are needed .
Approach: They explore the role of personas in the creation of emotional support conversations by LLMs.
Outcome: The proposed model can infer persona traits and maintain key persona characteristics while revealing shifts in emotionality and extraversion.
Empathy in Diversity: Personalized Depression and Anxiety Therapy via Dialogue State Tracking and Patient-Aware Planning (2026.acl-long)

Copied to clipboard

Challenge: Recent efforts have turned to large language models (LLMs) as therapeutic agents for psychological therapy tasks, yet robustness across diverse patients remains underexplored.
Approach: They propose a realistic role-play protocol for evaluating therapeutic dialogue agents and a de-identified, expert-annotated corpus of therapist–patient dialogues.
Outcome: The proposed framework outperforms baselines on therapeutic outcomes and dialogue quality while improving conversational efficiency.
LLM-guided Plan and Retrieval: A Strategic Alignment for Interpretable User Satisfaction Estimation in Dialogue (2025.naacl-long)

Copied to clipboard

Challenge: Existing methods for estimating user satisfaction with dialogue systems face challenges due to limited understanding of underlying reasons for user dissatisfaction and high costs of annotating user intentions.
Approach: They propose an interpretable framework for effective user satisfaction prediction . they propose to align utterances with strategies and large language models to retrieve relevant features from utterations.
Outcome: The proposed framework achieves state-of-the-art performance on three benchmarks for the USE task.
DischargeSim: A Simulation Benchmark for Educational Doctor–Patient Communication at Discharge (2025.emnlp-main)

Copied to clipboard

Challenge: Discharge communication is a critical yet underexplored component of patient care, where the goal shifts from diagnosis to education.
Approach: They propose a benchmark that evaluates large language models’ ability to act as personalized discharge educators.
Outcome: Experiments with 18 LLMs show that model size does not always yield better education outcomes, highlighting trade-offs in strategy use and content prioritization.
Exploring Persona Sentiment Sensitivity in Personalized Dialogue Generation (2025.acl-long)

Copied to clipboard

Challenge: Personalized dialogue systems have advanced with the integration of user-specific personas into large language models (LLMs).
Approach: They propose a dialogue generation approach that explicitly accounts for persona polarity by combining a turn-based generation strategy with a profile ordering mechanism and sentiment-aware prompting.
Outcome: The proposed approach accounts for persona polarity by combining a turn-based generation strategy with a profile ordering mechanism and sentiment-aware prompting.
3MDBench: Medical Multimodal Multi-agent Dialogue Benchmark (2025.emnlp-main)

Copied to clipboard

Challenge: Large Vision-Language Models (LVLMs) are being explored in medicine but their ability to conduct complex real-world telemedicine consultations remains underexplored.
Approach: They propose to use large vision-language models to conduct telemedicine consultations using a framework that simulates patient variability and evaluates diagnostic accuracy and dialogue quality via Assessor Agent.
Outcome: The proposed framework compares diagnostic strategies for open and closed-source LVLMs and shows that multimodal dialogue improves F1 score by 6.5% over non-dialogue settings.
Enhancing Goal-oriented Proactive Dialogue Systems via Dynamic Multi-dimensional Consistency Optimization (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing work on goal-oriented proactive dialogue systems failed to address the multi-dimensional consistency issue between generated responses and key contextual elements.
Approach: They propose a Dynamic Multi-dimensional Consistency Reinforcement Learning framework which measures the impact of each consistency dimension on overall dialogue quality and provides feedback to improve response quality.
Outcome: The proposed framework significantly improves the consistency of generated responses on two datasets.
AV-Dialog: Spoken Dialogue Models with Audio-Visual Input (2026.acl-long)

Copied to clipboard

Challenge: AV-Dialog uses audio and visual cues to track the target speaker, predict turn-taking, and generate coherent responses.
Approach: They propose a multimodal dialog framework that uses both audio and visual cues to track the target speaker.
Outcome: AV-Dialog outperforms audio-only models under interference, reducing transcription errors, improving turn-taking prediction and human-rated dialogue quality.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations